Papers with alternating training

1 papers
Guided Dialogue Policy Learning without Adversarial Learning in the Loop (2020.findings-emnlp)

Copied to clipboard

Challenge: Reinforcement learning methods suffer from sparse and unstable reward signals . alternating training of dialogue agent and reward model can get stuck in local optima .
Approach: They propose to decompose adversarial training into two steps to improve dialogue policy learning.
Outcome: The proposed method achieves remarkable task success rate using both on-policy and off-poly reinforcement learning methods.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations